Features

Pricing Molecules with AI: What Molecular Structure Can—and Cannot—Predict

Can the price of pharmaceutical molecules be predicted by artificial intelligence alone from molecular structure?

AI-driven molecular analysis is being explored to predict pharmaceutical intermediate pricing. (Photo credit: stock.adobe.com/angellodeco)

Editor’s Take: AI can extract a meaningful pricing signal from molecular structure, but commercial factors still account for most of the variation.

Recent advances in cheminformatics and artificial intelligence (AI) have enabled increasingly accurate predictions of molecular properties, synthetic accessibility, and feasible synthetic routes directly from molecular structure.

Whether similar approaches can be extended to predict the commercial value of organic intermediates remains an open question.

To investigate this hypothesis, we have developed and trained a multi-layer machine-learning architecture using more than 7,000 pharmaceutical intermediates represented through SMILES strings and their market price.

The objective was not just to build a pricing engine but more fundamentally to evaluate how much pricing information is encoded in molecular structure and therefore extractable through modern AI techniques.

The model recovers approximately one-third of the observed price variability from just the molecular structure, demonstrating that molecular structure contains a measurable—but incomplete—economic signal that can be extracted using the machine-learning methods applied in this study.

These results highlight the potential but also the limitations of structure-based pricing models.

Introduction

Over the last two decades, breakthroughs have been achieved in applying cheminformatics and artificial intelligence tools to molecular design and process development.

Examples include computer-assisted synthesis planning platforms like ASKCOS or Synthia, and molecular scoring systems such as Reaxys, SAS, SCScore and FSC. These tools estimate synthetic accessibility, route complexity and the synthesis feasibility of target molecules based on their structure.

While such approaches provide valuable information on how a molecule can be synthesized, these do not directly address the key question: Can the price of an organic intermediate be inferred from its molecular structure?

Prima facie, the hypothesis seems plausible because molecular structure influences synthesis complexity, stereochemical requirements, functional group density, precursor availability, and ultimately production costs.

However, contrary to melting point, toxicity, or solubility—price is not an intrinsic molecular property. Rather, it results from the interaction between chemistry, technology, supply chains, and competitive dynamics.

The objective of our work has therefore been to investigate the extent to which price information is encoded in molecular structure and whether modern AI methods can successfully extract and quantify this signal.

Chemistry, complexity and price

Tools like ASKCOS—an open-source software suite for computer-aided synthesis planning developed by MIT and E. Merck’s Chematica, and now called Synthia—allow chemists to rapidly identify feasible synthetic routes and their complexity to access target molecules whose structure has been translated into computer-readable form, such as SMILES strings or Molfiles.

Similarly, multiple systems like SAS (Synthesis Accessibility Score), SCS (Synthesis Complexity Score), or FSC (Focused Synthesizability Score) based on Machine Learning trained on millions of different reactions are routinely used in early-stage drug discovery to weed out molecules with high theoretical scores, a proxy for difficult, if not impossible, synthesis.

However, while these tools are useful in assessing how a molecule can be synthesized by providing a quantitative view of its synthesis complexity, they do not say what it will cost—synthesis complexity and price being related but not identical concepts.

Machine-learning approaches for predicting the price of chemical compounds have so far followed three main directions.

The first relies on graph neural networks trained directly on large commercial catalogs, learning the relationship between molecular structure and listed prices.1 While this approach has the advantage of covering very large datasets—the representativeness of catalog prices is liable to be challenged. Consequently, these models learn a combination of structural and commercial effects rather than the contribution of molecular structure alone.

A second approach estimates price from predicted synthetic routes combined with the cost of starting materials.2 These methods can potentially achieve higher accuracy when reliable retrosynthetic information is available – but are dependent on the quality of the route prediction itself and cannot be considered structure-only approaches

A third direction derives accessibility or cost-related scores from market data using contrastive or self-supervised learning techniques.3 These methods are mainly intended to rank compounds according to expected cost or accessibility rather than predict an absolute market price

A molecule can be synthetically challenging yet inexpensive if produced at a large scale through an optimized process. Conversely, a structurally simple molecule may have a high price due to puny volumes, scarce feedstock availability, or stringent containment requirements due to its potency.

The development of AI models trained on large product datasets and market prices to estimate the price of organic intermediates directly from molecular structure would therefore represent a compelling innovation and address substantial unmet needs.

Such models extending the structure-based analysis beyond synthetic feasibility, combining it with a price prediction, would have the potential to provide early in the product development process data-driven guidance on cost determinants, allowing screening for alternative intermediates or synthesis pathways. Expected benefits would include supporting informed decision-making across R&D and procurement, providing an integrated, predictive tool that bridges chemistry and economics.

To this end, we have combined the multi-year experience accumulated in the development and sourcing of organic intermediates for the pharmaceutical industry with expertise in AI and MLM (Machine Learning Models) for building a model to test the feasibility of predicting the price of new intermediates based solely on their molecular structure and assessing how much of the price can be traced to this single variable.

Experimental approach

Dataset

The input data comprise more than seven thousand commercially relevant pharmaceutical intermediates. For each compound, the molecular structure was represented by its SMILES string together with the latest market price; the CAS number was used for compound identification where available.

To reduce variability associated with geography, exchange rates, and transaction scale, all prices were normalized using quotations from Chinese producers for volumes of 1,000 kg and applying 2026 RMB/USD exchange rates.

The key descriptors of the molecule are provided by CAS/SMILES in a machine-readable format “understandable” by neural network models. At the same time, the price is the target variable that the model aims to learn and predict, linking it to the structural features of the molecule.

Although a dataset of this size is significant by industrial standards, it represents only a small fraction of the broader chemical universe. Consequently, some prediction errors are inevitably attributable to an incomplete coverage of chemical space. However, as discussed later, several systematic behaviors observed in the model suggest that limited dataset size alone does not fully explain the observed performance limits.

Model architecture

The architecture of the model developed is schematically illustrated in Figure 1.


Figure 1. AI-Based Molecular Price Prediction Model Architecture


It consists of a multilayered machine learning system designed to estimate the price of organic pharmaceutical intermediates in USD/kg directly from their molecular structure. It combines traditional cheminformatics, modern machine learning techniques, and large language model (LLM)–derived insights into a single predictive framework.

Features

The first step consists of converting the molecular structure into a set of numerical descriptors that machine-learning algorithms can process. These descriptors include well-established chemical characteristics such as molecular weight, polarity, ring structures, and other features reflecting the molecule composition and complexity.

In parallel, additional descriptors are generated using a Large Language Model (LLM) to capture higher-level chemical information, like perceived synthetic complexity and key functional groups.

As this process generates multiple variables, only those providing meaningful information are retained. The selected descriptors can also be combined to create new variables, helping to identify more complex relationships between molecular structure and price.

The dataset is then divided into separate subsets used for training, validation, and testing. To ensure that the results are robust and not dependent on a particular data split, the model is trained and evaluated several times using different data subsets, which also reduces overfitting risk.

Model training

The dataset was divided into independent training, validation, and test subsets.

Model development, parameter optimization, and architecture selection were performed exclusively on the training and validation data.

Final performance was evaluated on a held-out test set comprising molecules never seen during model training. Cross-validation procedures were additionally applied to reduce dependence on any specific train-test split and provide a more robust assessment of predictive performance.

Model architecture

The model combines several complementary artificial intelligence approaches, each designed to identify different relationships between molecular structure and price. Some models focus on simple and easily interpretable relationships, while others capture more complex patterns that may not be immediately apparent.

In addition, specialized models are trained to perform particularly well on specific families of molecules or regions of chemical space. This allows the system to adapt its behavior depending on the type of compound analyzed.

The predictions generated by the various models are subsequently combined into a single estimate. Additional refinements include comparing the molecule under consideration with structurally similar compounds in the training dataset—the hypothesis being that similar molecules often exhibit comparable pricing patterns.

A final AI-based review step is then applied to identify unusual cases and improve predictions for rare or particularly complex molecular structures.

Ultimately, the multilayered predictive system involving multiple models working together receives the molecular structure of a compound as input and generates the estimated market price as output.

5. Results

Table 1 summarizes dataset size and predictive performance.



5.1 General predictive performance

The model’s ability to assess molecules that it had not been exposed to has been evaluated through an independent test set comprising 869 pharmaceutical intermediates.

The overall performance achieved an R² value of 0.36 on the price logarithm, corresponding to a Root Mean Square Error (RMSE) of 0.56 and a Mean Absolute Error (MAE) of 0.44 (see Figure 2).


Figure 2. Predicted vs. Actual Pharmaceutical Intermediate Prices in the Holdout Test Set


The R² value of 0.36 indicates that approximately one-third of the observed price variability can be recovered from just the molecular structure. The remaining variability derives from non-structural price determinants and/or structural information that the model does not fully capture.

These findings suggest that based solely on the molecular structure, the model can explain approximately one-third of the observed variation in market prices – confirming that molecular structure provides a measurable economic signal and that machine-learning techniques can partly link chemistry with commercial value.

However, the practical implications of these metrics require careful interpretation. The typical prediction error for an individual molecule remains approximately a factor of 2.3, with only 43% of compounds predicted within a factor of two of their actual market prices.

Therefore, while the model successfully distinguishes broad pricing regions within the chemical space, its ability to predict the absolute prices of specific products is limited, confirming that molecular structure alone provides incomplete information on pricing.

5.2 Systematic prediction compression

Analysis of the predicted-versus-observed prices reveals a clear systematic behavior. Predictions are compressed towards the center of the distribution, with a best-fit slope of approximately 0.32 compared with the ideal value of 1.0.

In practical terms, the model tends to overestimate the cheapest molecules and underestimate the most expensive ones:

• The lowest-priced 20% of compounds being predicted to have a price of approximately 64 USD/kg versus an average actual market price of around 16 USD/kg.

• The highest-priced 20% having an average market price in the order of 2,231/USD/kg to be compared to the 339 USD/kg predicted by the model.

The observed compression suggests that molecular structure captures broad determinants of cost but not the drivers responsible for extreme prices.

5.3 Performance after removal of price extremes

To evaluate the influence of extreme values, molecules priced below USD 20/kg and above USD 20,000/kg were removed from the test set.

The resulting improvement was limited—the median prediction error decreasing from approximately 2.3-fold to 2.1-fold, while the proportion of compounds predicted within a factor of two increased from 43% to 48%.

The modest improvement obtained after removing outliers suggests that the model’s limitations are not confined to a few exceptional compounds but are instead linked to broader structural characteristics.

5.4 What the model reveals about chemical price formation

Taken together, these results support two important conclusions.

First, molecular structure undeniably contains information relevant to pricing. Functional groups, stereochemistry, molecular complexity and other structural features influence synthetic accessibility and manufacturing costs. The model successfully extracts part of this information, using it to explain approximately one-third of the observed price variation.

Second, molecular structure alone does not explain most market price differences.

Unlike physicochemical properties such as molecular weight, boiling point, density, solubility, or toxicity, price is not an intrinsic molecular property. Rather, it results from the interaction between molecular structure, manufacturing technology, production scale, supply-chain structure, and competitive dynamics. This distinction helps explain why predicting price from structure alone is inherently more difficult than predicting traditional molecular properties.

Conclusion

Our study suggests that about one-third of the variability in pharmaceutical intermediate prices can be inferred directly from the molecular structure using the machine-learning techniques applied.

Although the resulting predictions are insufficient for accurate commercial pricing, these confirm that molecular structure contains a measurable economic signal.

Substantial improvement of predictive performance can be expected using future models combining structural descriptors with synthetic processes and commercial information.

References

1. Sanchez-Garcia, R.; Havasi, D.; Takács, G.; Robinson, M. C.; Lee, A.; von Delft, F.; Deane, C. M. CoPriNet: graph neural networks provide accurate and rapid compound price prediction for molecule prioritization. Digital Discovery 2023, 2, 103–111. doi:10.1039/D2DD00071G.

2. Abderrahmane, M.; Tajmouati, H.; Barros Ribeiro da Silva, V.; Perron, Q. Predicting the price of molecules using their predicted synthetic pathways. Molecular Informatics 2025, 44 (2), 202400039. doi:10.1002/minf.202400039.

3. Hastedt, F.; Hellgardt, K.; Yaliraki, S.; et al. MolPrice: assessing synthetic accessibility of molecules based on market value. Journal of Cheminformatics 2025, 17, 150. doi:10.1186/s13321-025-01076-3.


Dr. Michele Jermini is the Managing Director of Exeris ([email protected]).

Dr. Paul Hanselmann is Founder and CEO of ChemSynthDesign GmbH.

Federico Brooks is Founder and CTO of Custodian Privacy Sagl.

Dr. Enrico Polastro is a Vice-President of Arthur D.Little.


Keep Up With Our Content. Subscribe To Contract Pharma Newsletters